Remote sensing image super-resolution method and system based on hierarchical progressive distillation
Patent Information
- Application Number
- CN202610867698.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
本申请提供的基于分层渐进蒸馏的遥感图像超分辨率方法及系统,通过浅层特征提取单元对输入图像进行初步特征建模,实现对基础纹理信息的有效表达;通过分层渐进蒸馏单元对特征进行多阶段逐级筛选与压缩,在减少冗余信息的同时保留关键结构特征,从而提高特征表达效率;通过空间-频域混合注意力增强块对特征进行空间依赖建模与频域信息补充,增强不同尺度下的细节响应能力;在此基础上,通过高分辨率图像重建单元实现高分辨率图像重建,在保持较低模型复杂度与计算开销的同时,提高遥感图像的细节恢复能力与结构保持能力。
Smart Images

Figure CN122820439A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual remote sensing image technology, and more specifically to a method and system for super-resolution of remote sensing images based on hierarchical progressive distillation. Background Technology
[0002] In recent years, deep learning-based image super-resolution (SR) reconstruction methods have been widely applied in the field of remote sensing image processing. Thanks to the powerful modeling capabilities of convolutional neural networks for multi-scale features, these methods have made significant progress in improving spatial resolution and detail recovery. However, remote sensing images typically feature high resolution, wide coverage, and large data volumes, leading to high computational complexity and storage overhead for existing hierarchical progressive distillation networks in practical applications. This dependence on computational resources limits their deployment and application on resource-constrained platforms such as edge computing devices and embedded systems.
[0003] To address the aforementioned issues, researchers have proposed various optimization strategies, including lightweight network design, feature reuse, and knowledge distillation, to reduce model complexity and improve operational efficiency. Among these, the distillation-based approach progressively filters and compresses feature information, minimizing redundant computation while preserving reconstruction performance as much as possible. However, during feature compression, some crucial information may be lost, affecting the final reconstruction quality. Furthermore, remote sensing images typically contain rich textural details and complex spatial structures, with significant differences between targets at different scales, rendering existing methods insufficient in balancing global semantic information with local detail restoration.
[0004] Compared to super-resolution tasks for natural images, super-resolution of remote sensing images focuses more on restoring structural information of ground features, such as building edges, road networks, and surface textures. This structural information often exhibits distinct spatial distribution characteristics and frequency properties in images, making it prone to blurring or distortion under low-resolution conditions. Therefore, relying solely on single spatial domain features for modeling is insufficient to fully characterize the complex information in remote sensing images, limiting the model's reconstruction capabilities.
[0005] Existing methods typically optimize the model based on pixel-level reconstruction errors during training. While these methods offer advantages in overall reconstruction accuracy, they still fall short in recovering high-frequency details and structural information. Meanwhile, improving feature representation capabilities and information utilization efficiency while maintaining model lightweightness has become a critical issue that urgently needs to be addressed in the field of remote sensing image super-resolution. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a method and system for super-resolution of remote sensing images based on hierarchical progressive distillation. This method can effectively screen and enhance features at multiple stages, and also takes into account the ability to model information in both the spatial and frequency domains, thereby achieving super-resolution of remote sensing images. This is of great significance for improving reconstruction quality and practical application performance.
[0007] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0008] According to a first aspect of this application, a remote sensing image super-resolution method based on hierarchical progressive distillation is provided, comprising: Acquire raw remote sensing images and process them to obtain pseudo-high resolution images; The pseudo-high-resolution image is input into a pre-constructed hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high-resolution image reconstruction unit. Based on the shallow feature extraction unit, the pseudo-high resolution image is encoded to obtain shallow features; In each of the hierarchical progressive distillation units, the shallow features are subjected to multi-stage feature processing. During the step-by-step processing, feature selection and information compression are performed to obtain initial deep features. Furthermore, the initial deep features are subjected to spatial modeling and frequency domain enhancement processing to obtain enhanced features. Based on the global feature fusion unit, the enhanced features output by each hierarchical progressive distillation unit are spliced and fused, and then fused with the shallow features through residual connections to obtain deep fused features.
[0009] In some embodiments of this application, based on the foregoing scheme, encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features includes: A lightweight feature extraction method is used to encode the input image to obtain shallow features.
[0010] In some embodiments of this application, based on the aforementioned scheme, the hierarchical progressive distillation unit includes a first hierarchical progressive distillation block, a second hierarchical progressive distillation block, and a third hierarchical progressive distillation block, performing multi-stage feature processing on the shallow features. During the step-by-step processing, feature filtering and information compression are performed to obtain initial deep features, including: The shallow features are input into the first hierarchical progressive distillation block, and the shallow features are processed based on the first channel attention to obtain the first distillation features. The shallow features are then processed based on the first partial large kernel convolution to obtain the first refined features. The first distillation feature and the first refining feature are input into the second hierarchical progressive distillation block. The first refining feature is processed based on the second channel attention to obtain the second distillation feature. The first refining feature is processed based on the second partial large kernel convolution to obtain the second refining feature. The second distillation feature and the second refining feature are input into the third hierarchical progressive distillation block. The second refining feature is processed based on the third channel attention to obtain the third distillation feature. The second refining feature is processed based on the third partial large kernel convolution to obtain the third refining feature. The fourth refining feature is obtained by processing the third refining feature through a single BSConv operation. The initial deep characteristics are obtained by fusing the first distillation characteristics, the second distillation characteristics, the third distillation characteristics, and the fourth refining characteristics.
[0011] In some embodiments of this application, based on the foregoing scheme, the hierarchical progressive distillation unit further includes a spatial-frequency domain hybrid attention enhancement block, which includes a dynamic window attention subunit and a multi-scale frequency domain subunit. The step of performing spatial modeling and frequency domain enhancement processing on the initial deep features to obtain enhanced features includes: The initial deep features are input into the dynamic window attention subunit to obtain spatially enhanced features; The initial deep features are input into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features; The initial deep features, the spatial enhancement features, and the frequency domain enhancement features are self-learned and fused to obtain the initial enhancement features; The initial enhanced features are sequentially processed by convolutional projection pixel normalization, and then added to the shallow features through residual connections to obtain the final enhanced features.
[0012] In some embodiments of this application, based on the foregoing scheme, the step of inputting the initial deep features into a dynamic window attention subunit to obtain spatially enhanced features includes: The optimal window size for the shallow features is obtained based on a lightweight prediction network. Based on the optimal window size, the initial deep features are divided into several windows. Multi-head self-attention is calculated in each window to obtain spatially enhanced features.
[0013] In some embodiments of this application, based on the foregoing scheme, the step of inputting the initial deep features into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features includes: The first frequency domain feature corresponding to the initial deep feature is obtained based on the Fast Fourier Transform; The first frequency domain feature is subjected to downsampling, feature enhancement and upsampling operations at a predefined scale to obtain the second frequency domain feature corresponding to the predefined scale. After fusing the second frequency domain features corresponding to each scale, adding them to the first frequency domain features, and then mapping them back to the spatial domain through inverse fast Fourier transform, a spatial attention map is generated. The initial deep features are processed based on the spatial attention map to obtain frequency domain enhanced features.
[0014] In some embodiments of this application, based on the foregoing scheme, the step of splicing and fusing the enhanced features output by each of the hierarchical progressive distillation units based on the global feature fusion unit, and then fusing them with the shallow features through residual connections to obtain deep fused features, includes: The enhancement features are gradually refined through several hierarchical progressive distillation blocks to obtain several refined features; Several refined features are spliced and fused together to obtain deep fused features; The shallow features and the deep fusion features are aggregated through residual connections to obtain deep fusion features, which are then reconstructed into a high-resolution remote sensing image through pixel rearrangement operations.
[0015] According to a second aspect of this application, a remote sensing image super-resolution system based on hierarchical progressive distillation is provided, comprising: The first acquisition module is used to acquire raw remote sensing images and process the raw remote sensing images to obtain pseudo-high resolution images. The model building module is used to input the pseudo-high-resolution image into a pre-built hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high-resolution image reconstruction unit. The second acquisition module is used to encode the pseudo-high resolution image based on the shallow feature extraction unit to obtain shallow features; The third acquisition module is used to perform multi-stage feature processing on the shallow features in each of the hierarchical progressive distillation units, and to perform feature screening and information compression in the step-by-step processing to obtain initial deep features, and to perform spatial modeling and frequency domain enhancement processing on the initial deep features to obtain enhanced features. The fourth acquisition module is used to perform splicing and fusion processing on the enhanced features output by each of the hierarchical progressive distillation units based on the global feature fusion unit, and then fuse them with the shallow features through residual connection to obtain deep fusion features; The fifth acquisition module is used to perform convolution and pixel rearrangement processing on the deep fusion features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
[0016] According to a third aspect of this application, a computer-readable storage medium is provided that stores a computer program thereon, the computer program including executable instructions that, when executed by a processor, implement the method described above.
[0017] According to a fourth aspect of this application, an electronic device is provided, comprising: One or more processors; A memory for storing executable instructions of the processor, which, when executed by the one or more processors, cause the one or more processors to implement the method described above.
[0018] The beneficial effects of this application are as follows: The remote sensing image super-resolution method and system based on hierarchical progressive distillation provided in this application performs preliminary feature modeling on the input image through a shallow feature extraction unit, achieving effective expression of basic texture information; the hierarchical progressive distillation unit performs multi-stage progressive screening and compression of features, reducing redundant information while retaining key structural features, thereby improving feature expression efficiency; a spatial-frequency domain hybrid attention enhancement block performs spatial dependency modeling and frequency domain information supplementation on features, enhancing the detail response capability at different scales; on this basis, a high-resolution image reconstruction unit achieves high-resolution image reconstruction, improving the detail recovery and structure preservation capabilities of remote sensing images while maintaining low model complexity and computational overhead.
[0019] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and are intended to explain the invention, but do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart of a remote sensing image super-resolution method based on hierarchical progressive distillation according to the present invention; Figure 2 This is a hierarchical progressive distillation network diagram of the remote sensing image super-resolution method provided by the present invention; Figure 3 This is a detailed architecture diagram of the hierarchical progressive distillation unit of the remote sensing image super-resolution method provided by the present invention; Figure 4 This is a block diagram of the dynamic window attention mechanism of the remote sensing image super-resolution method provided by the present invention; Figure 5 This is a block diagram of the multi-scale frequency domain attention mechanism of the remote sensing image super-resolution method provided by the present invention; Figure 6 This is a comparison chart of the first subjective visual effects of different models of this invention on the AID dataset; Figure 7 This is a comparison of the first subjective visual effects of different models of this invention on the RSCN7 dataset; Figure 8 This is a schematic diagram of a remote sensing image super-resolution system based on hierarchical progressive distillation according to the present invention. Figure 9 This is a schematic diagram of an electronic device according to the present invention. Detailed Implementation
[0021] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0022] It should be understood that the terms "comprising" and other similar expressions in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, or apparatus that includes a series of steps or units and is not limited to the listed steps or units. Furthermore, "first" and "second" are used to distinguish different objects and are not intended to describe a specific order.
[0023] According to the first aspect of this application, Figure 1 As shown, this embodiment provides a remote sensing image super-resolution method based on hierarchical progressive distillation, including: Step S1: Acquire the original remote sensing image and process the original remote sensing image to obtain a pseudo-high resolution image.
[0024] In some embodiments of this example, an upsampling unit is set up, and a bicubic interpolation method is used to upsample low-resolution remote sensing images. Upsampled to a target high-resolution image benchmark With the same spatial resolution, a pseudo-high resolution image is obtained. ,in Indicates the image height. Indicates the image width. Indicates the number of channels. Represents the real number field.
[0025] In one specific embodiment of this application, the publicly available remote sensing image datasets AID and RSCN7 are preferably used. The AID (Aerial Image Dataset) dataset is a remote sensing scene image dataset compiled and constructed by institutions such as Huazhong University of Science and Technology. It is mainly used for remote sensing scene classification and understanding tasks. The data comes from publicly available remote sensing image platforms such as Google Earth and has high spatial resolution and rich information on land cover categories. The dataset contains various typical remote sensing scene categories, such as airports, commercial areas, industrial areas, residential areas, forests, and rivers, exhibiting strong scene complexity and diversity, and effectively reflecting the spatial distribution characteristics of different land cover structures in remote sensing images. The AID dataset contains 10,000 images, covering 30 scene categories, with each category containing approximately 200 to 420 images. The image resolution is 600×600 pixels. In this invention, the last 15 images in each category are selected as the test set, and the remaining images are used for training and validation. The RSSCN7 (Remote Sensing Scene Classification with Seven Categories) dataset is a publicly available dataset for remote sensing scene analysis, containing seven typical remote sensing scene categories, such as forests, coastlines, residential areas, industrial areas, fields, rivers, and farmland. The images in this dataset were collected from remote sensing images taken at different times, seasons, and under different imaging conditions, exhibiting significant scale and illumination variations, making them suitable for evaluating the robustness and generalization ability of models in complex environments. The dataset consists of 2,800 images, divided into seven categories, with each category containing 400 images at a resolution of 400×400. In this embodiment, the last 40 images from each category are used as the test set, with the remainder used for training and validation. During data processing, all training, validation, and test images are first downsampled to 256×256 resolution using bicubic interpolation as high-resolution (HR) reference images, and then further downsampled at 2x and 4x scales to generate corresponding low-resolution (LR) images.
[0026] Step S2: Input the pseudo-high resolution image into a pre-constructed hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high resolution image reconstruction unit.
[0027] In some embodiments of this example, the hierarchical progressive distillation network is trained using a training set to obtain a trained hierarchical progressive distillation network, and pseudo-high resolution images are processed based on the trained hierarchical progressive distillation network.
[0028] Step S3: Based on the shallow feature extraction unit, the pseudo-high resolution image is encoded to obtain shallow features.
[0029] In some embodiments of this example, encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features includes: The input image is encoded using a lightweight feature extraction method to obtain shallow features. .
[0030] In one specific embodiment, the input pseudo-high-resolution image Basic feature extraction is performed, and preliminary modeling of low-level texture and local structure information is achieved through lightweight convolution operations and channel expansion, outputting shallow features. .
[0031] like Figure 2 As shown, the shallow feature extraction unit, as the first-stage component of the hierarchical progressive distillation network, provides basic feature inputs for subsequent hierarchical progressive distillation. The shallow feature extraction unit does not involve complex multi-branch structures; instead, it employs a lightweight feature extraction method to encode the input image, enabling the network to achieve stable shallow representation capabilities with low computational overhead. In a specific embodiment of this application, the shallow feature extraction process can be represented as follows: ; in, Indicates shallow features. This means copying the original input image along the channel dimension. The result obtained after that This indicates a channel-level concatenation operation. This indicates a lightweight convolutional feature extraction operation.
[0032] By using shallow feature extraction units to initially encode the input image, the hierarchical progressive distillation network can reduce the computational complexity of subsequent feature extraction stages while preserving basic spatial structure information. This provides a more stable and semantically grounded shallow feature representation for hierarchical progressive distillation. .
[0033] Through this shallow feature extraction process, the hierarchical progressive distillation network can effectively preserve the basic texture information and local structural information in remote sensing images, providing the necessary feature foundation for subsequent hierarchical progressive distillation's step-by-step feature selection and information compression, thereby improving the overall network's feature expression ability and reconstruction performance.
[0034] Step S4: In each of the layered progressive distillation units, the shallow features are subjected to multi-stage feature processing. During the step-by-step processing, feature selection and information compression are performed to obtain initial deep features. The initial deep features are then subjected to spatial modeling and frequency domain enhancement processing to obtain enhanced features.
[0035] In some implementations of this embodiment, such as Figure 3 As shown, the hierarchical progressive distillation unit includes a first hierarchical progressive distillation block, a second hierarchical progressive distillation block, and a third hierarchical progressive distillation block. Based on the hierarchical progressive distillation unit, multi-stage feature processing is performed on the shallow features. During the step-by-step processing, feature selection and information compression, as well as spatial modeling and frequency domain enhancement processing, are performed to obtain enhanced features, including: The shallow features are input into the first hierarchical progressive distillation block, and the shallow features are processed based on the first channel attention to obtain the first distillation features. The shallow features are then processed based on the first partial large kernel convolution to obtain the first refined features. The first distillation feature and the first refining feature are input into the second hierarchical progressive distillation block. The first refining feature is processed based on the second channel attention to obtain the second distillation feature. The first refining feature is processed based on the second partial large kernel convolution to obtain the second refining feature. The second distillation feature and the second refining feature are input into the third hierarchical progressive distillation block. The second refining feature is processed based on the third channel attention to obtain the third distillation feature. The second refining feature is processed based on the third partial large kernel convolution to obtain the third refining feature. The fourth refining feature is obtained by processing the third refining feature through a single BSConv operation. The initial deep characteristics are obtained by fusing the first distillation characteristics, the second distillation characteristics, the third distillation characteristics, and the fourth refining characteristics.
[0036] In this embodiment, the first, second, and third layered progressive distillation blocks not only acquire distillation features for progressive channel modeling and information filtering of shallow features, but also acquire refined features for progressive compression and reconstruction of features, thereby achieving progressive extraction of shallow features into initial deep features.
[0037] shallow features The input is fed into a three-stage cascaded distillation unit, and the first distillation characteristics are obtained in the first stage. With the first refining feature The second distillation characteristics are obtained in the second stage. Characteristics of the third distillation The third distillation characteristics are obtained in the third stage. Characteristics of the third distillation Finally, the fourth refining characteristic was obtained. This completes the multi-stage, step-by-step information filtering and feature compression process. The specific distillation process is as follows: In the first distillation stage, the most important features are highlighted by calculating the first channel attention: ; in, Indicates shallow features. and These are the learnable parameters for the first stage. This indicates the compression ratio. GELU (Gaussian Error Linear Unit) is an activation function, and GAP (Global Average Pooling) is global average pooling. This represents the Sigmoid activation function. The distillation characteristics of this stage are obtained through element-wise multiplication: ; in, This represents element-wise multiplication; In the feature refinement path, a Partial Large Kernel Convolution (PLKC) module is introduced to process shallow features: ; Among them, PLKC1 (Partial Large Kernel Convolution module 1) is the first part of large kernel convolution.
[0038] In the second stage, the distillation and refining processes continue in a hierarchical manner, with the first refining characteristics obtained in the first stage being... It also serves as input for both the distillation branch and the further refining branch: ; ; ; Among them, PLKC2 (Partial Large Kernel Convolution module 2) is the second part of large kernel convolution. and These are the learnable parameters for the second stage. 2 indicates the compression ratio.
[0039] The third stage completes the entire stratified distillation process: ; ; ; Among them, PLKC3 (Partial Large Kernel Convolution module 3) is the third part of large kernel convolution. and These are the learnable parameters for the third stage. 3 indicates the compression ratio.
[0040] At the end of the feature refinement path, the third refined feature is compressed and further optimized through a single BSConv operation: ; in, This is the fourth refining characteristic.
[0041] The three-stage hierarchical design of this embodiment ensures that features are progressively distilled at increasingly higher levels of abstraction, with each stage built upon the refined features of the previous stage. This hierarchical distillation mechanism enables the network to simultaneously capture low-level texture details and high-level semantic information at different stages.
[0042] The characteristics obtained from each stage of distillation are combined with the characteristics of the fourth refining process, and then... Convolutional processes are used to fuse the data, resulting in initial deep features. The calculation formula is as follows: ; in, This indicates a lightweight convolutional feature extraction operation.
[0043] In some embodiments of this example, the hierarchical progressive distillation unit further includes a spatial-frequency domain hybrid attention enhancement block, which includes a dynamic window attention subunit and a multi-scale frequency domain subunit. The process of performing spatial modeling and frequency domain enhancement on the initial deep features to obtain enhanced features includes: The initial deep features are input into the dynamic window attention subunit to obtain spatially enhanced features; The initial deep features are input into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features; The initial deep features, the spatial enhancement features, and the frequency domain enhancement features are self-learned and fused to obtain the initial enhancement features; The initial enhanced features are sequentially processed by convolutional projection pixel normalization, and then added to the shallow features through residual connections to obtain the final enhanced features.
[0044] In some implementations of this embodiment, such as Figure 4As shown, the step of inputting the initial deep features into a dynamic window attention subunit to obtain spatially enhanced features includes: The optimal window size for the shallow features is obtained based on a lightweight prediction network. Based on the optimal window size, the initial deep features are divided into several windows. Multi-head self-attention is calculated in each window to obtain spatially enhanced features.
[0045] Specifically, dynamic window attention aims to adaptively capture local spatial dependencies to overcome the limitations of fixed window partitioning. First, given initial deep features... The optimal window size is dynamically determined by analyzing its content features through a lightweight prediction network.
[0046] in, For convolution operations, It is the Sigmoid activation function. This represents element-wise multiplication. This represents the predicted window size. Based on this prediction, the feature map is divided into multiple windows, and multi-head self-attention is computed within each window: ; in, This represents the relative position offset obtained through a lookup table. , , These are the query vector, key vector, and value vector, respectively. Let be the dimension of the key vector. For normalization function, This indicates the transpose. This employs multi-head self-attention. After completing the attention calculation, the outputs of all windows are reassembled into a complete feature map. This dynamic partitioning strategy enables the network to automatically adjust the receptive field size based on the complexity of the image content: smaller windows are used in textured regions to capture details, while larger windows are used in flat regions to improve computational efficiency.
[0047] In some implementations of this embodiment, such as Figure 5 As shown, the step of inputting the initial deep features into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features includes: The first frequency domain feature corresponding to the initial deep feature is obtained based on the Fast Fourier Transform; The first frequency domain feature is subjected to downsampling, feature enhancement and upsampling operations at a predefined scale to obtain the second frequency domain feature corresponding to the predefined scale. After fusing the second frequency domain features corresponding to each scale, they are added to the first frequency domain features, and then mapped back to the spatial domain through inverse fast Fourier transform to generate frequency domain enhanced features. The initial deep features are processed based on the spatial attention map to obtain frequency domain enhanced features.
[0048] Specifically, multi-scale frequency domain attention enhances global features in the frequency domain by processing frequency components at different scales in parallel, thereby comprehensively capturing the structural information of the image. Given initial deep features... First, the first frequency domain features corresponding to the initial deep features are obtained through Fast Fourier Transform (FFT): ; in, For Fast Fourier Transform, and These are the real part features and the imaginary part features, respectively. The first frequency domain feature includes both real part features and imaginary part features.
[0049] In a predefined scale set The frequency domain enhancement features are processed in parallel. For each scale... The frequency domain enhancement features undergo downsampling, feature enhancement, and upsampling operations in sequence: ; ; in, This indicates that the feature is downsampled to its original size. , This indicates that the feature will be upsampled back to its original size. Indicates a three-layer A convolutional neural network (using the LeakyReLU activation function except for the last layer) is used to process frequency domain enhancement features after downsampling. and Representing the scale respectively Features of the real and imaginary parts after enhancement.
[0050] The enhanced features at each scale are fused and added to the first frequency domain features. Then, they are mapped back to the spatial domain using an inverse fast Fourier transform (IFFT) to generate a spatial attention map. ; ; ; in, Representing scale The second frequency domain feature consists of the real part and the imaginary part. constitute; Indicates the use of multi-scale features for fusing Convolution weight matrix; This represents the multi-scale frequency domain enhancement features after fusion; This represents the Sigmoid activation function; This represents element-wise multiplication; This represents the output features after multi-scale frequency domain attention enhancement. This is the inverse fast Fourier transform.
[0051] The mechanism of this invention enables the network to simultaneously focus on the global contour structure and local texture details of an image, achieving a true global receptive field in the frequency domain, while maintaining high computational efficiency through a multi-scale processing strategy.
[0052] In one specific embodiment of this application, the fused initial deep features Further enhancement via a spatial-frequency hybrid attention (SHA) block enables joint modeling of spatial and frequency domain information: ; in, For the initial enhancement features, and These are learnable parameters used to adaptively balance the contributions of frequency domain enhancement features and spatial features. Finally, the initial enhancement features are sequentially processed... The convolutional projection pixels are normalized and then added to the shallow features through residual connections to obtain the final enhanced features:
[0053] in, For the final enhanced features, This is a pixel normalization operation.
[0054] HPDB achieves comprehensive feature enhancement through its integrated architecture: it combines hierarchical distillation to preserve multi-level features, employs progressive refinement to improve feature quality, and introduces SHA to achieve comprehensive feature enhancement; at the same time, through careful design choices, it maintains computational efficiency while achieving the above functions.
[0055] Step S5: Based on the global feature fusion unit, the enhanced features output by each of the hierarchical progressive distillation units are spliced and fused, and then fused with the shallow features through residual connections to obtain deep fused features.
[0056] In some embodiments of this example, the step of splicing and fusing the enhanced features output by each of the hierarchical progressive distillation units based on the global feature fusion unit, and then fusing them with the shallow features through residual connections to obtain deep fused features, includes: The enhancement features are gradually refined through several hierarchical progressive distillation blocks to obtain several refined features; Several refined features are spliced and merged to obtain deeply fused features.
[0057] Specifically, A series of cascaded hierarchical progressive distillation blocks (HPDB) progressively refine the enhanced features:
[0058] in, For the first A refined feature.
[0059] After all the features of the intermediate blocks are spliced and merged, a rich set of deep fused features is formed:
[0060] in, This is a feature of deep fusion.
[0061] Step S6: Based on the high-resolution image reconstruction unit, perform convolution and pixel rearrangement processing on the deep fusion features to obtain a high-resolution remote sensing image.
[0062] In some embodiments of this example, high-resolution remote sensing images The calculation process is as follows:
[0063] in, This indicates a convolution operation with a kernel size of 3×3. This represents a pixel rearrangement operation used to achieve spatial reconstruction from feature maps to high-resolution images. This represents the final output high-resolution remote sensing image.
[0064] In one specific embodiment of this application, shallow features are used... With deep fusion features Element-wise addition is performed to effectively fuse low-frequency structural information with high-frequency detail information; then, convolution operation is used to perform channel mapping on the fused features; finally, pixel rearrangement operation is used to restore spatial resolution, thereby completing the reconstruction process of remote sensing image.
[0065] Thus, this embodiment will enhance the initial characteristics output by each layered progressive distillation unit. Perform channel-dimensional splicing, and sequentially pass through Feature fusion is performed using convolution, GELU activation function, and blueprint separating convolution to obtain fused features containing rich multi-scale representation information. The fused features are combined with the shallow features. Deep fusion features are obtained by element-wise addition through residual connections, and these features are then input into the high-resolution image reconstruction unit. High-resolution remote sensing images are reconstructed using convolution and pixel rearrangement operations. .
[0066] Furthermore, during the training of the hierarchical progressive distillation network, a pixel-level reconstruction loss function (such as the L1 loss function) is used to optimize the network parameters. By constraining the pixel differences between the reconstructed image and the real high-resolution image, the hierarchical progressive distillation network is guided to learn more accurate structural information and texture details, thereby improving the overall quality of super-resolution reconstruction of remote sensing images.
[0067] This application achieves the mapping from low-resolution remote sensing images to high-resolution images by performing forward inference on the trained hierarchical progressive distillation network. While ensuring that the computational complexity of the hierarchical progressive distillation network is controllable, it effectively improves the detail representation and structural consistency of the reconstructed images, thereby meeting the application requirements of fine reconstruction of remote sensing images.
[0068] In one specific embodiment of this application, the remote sensing image super-resolution testing process includes the following steps: The low-resolution original remote sensing image is denoted as Its image size is First, it is upsampled to obtain a pseudo-high resolution image. Its size is ,in , Indicates the image height. , Indicates the image width; pseudo-high resolution image Input the shallow feature extraction unit to obtain shallow features ; shallow features Inputting a hierarchical progressive distillation unit yields initial deep features. ; Initial deep features Input spatial-frequency domain hybrid attention enhancement blocks yield output enhancement features of hierarchical progressive distillation units. ; The intermediate enhancement features output from each stratified progressive distillation unit Perform channel-dimensional splicing, and sequentially pass through Feature fusion is performed using convolution, GELU activation function, and blueprint separating convolution to obtain fused features containing rich multi-scale representation information. The fused features are combined with the shallow features. Deep fusion features are obtained by element-wise addition through residual connections. These deep fusion features are then input into a high-resolution image reconstruction unit. High-resolution remote sensing images are reconstructed using convolution and pixel rearrangement operations. .
[0069] During the inference stage of the hierarchical progressive distillation network, low-resolution remote sensing images from the test data are... The trained hierarchical progressive distillation network is input, and the corresponding high-resolution reconstruction result is directly obtained through the aforementioned forward propagation process. When the output of the hierarchical progressive distillation network meets the preset evaluation criteria, the hierarchical progressive distillation network can be applied to actual remote sensing image super-resolution reconstruction tasks, thereby achieving high-quality reconstruction of low-resolution remote sensing images.
[0070] like Figure 6 As shown, a specific embodiment is used in this embodiment to illustrate the problem. The embodiment uses the AID dataset, which contains 10,000 images covering 30 scene categories. Each category contains approximately 200 to 420 images, and the image resolution is 600×600 pixels.
[0071] During the data processing phase, all training, validation, and test images were first downsampled to 256×256 resolution using bicubic interpolation as high-resolution (HR) reference images. These were then further downsampled at 2x and 4x scales to generate corresponding low-resolution (LR) images. In each category of the dataset, the last 15 images were selected as the test set, with the remaining images used for training and validation. This data setup covers different terrain structures, comprehensively considers training needs in various scenarios, and simulates the model's performance under diverse conditions.
[0072] The SR reconstruction results were evaluated using three metrics: peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and feature similarity, in order to verify the performance of SR reconstruction under the luminance channel.
[0073] During model training, the batch size was set to 32, the input image was randomly cropped into 32×32 patches, and common data augmentation strategies were employed, including random rotation and horizontal flipping. The optimizer used was the Adam optimizer, with parameters set to... , , .
[0074] In the ×2 super-resolution task, the model is trained from scratch for 500k iterations, with an initial learning rate of [value missing]. The learning rate scheduling employs a cosine annealing restart strategy, with a period set to [200000, 200000, 100000], a restart weight of [1, 0.5, 0.3], and a minimum learning rate set to... .
[0075] In the ×4 super-resolution task, fine-tuning is performed based on the pre-trained ×2 model, which is also trained for 500k iterations with the same initial learning rate. At the same time, a restart strategy with a period of [200000, 200000, 100000] and a weight of [1, 0.5, 0.3] is used.
[0076] All experiments were implemented using PyTorch and were trained and tested on two NVIDIA GeForce RTX 4090 GPUs.
[0077] The remote sensing image SR methods selected for comparison include: Bicubic, HAN, MHAN, CTNet, HSENet, MSGAN, LKFN, TTST, CATANet, and FMSR. Bicubic is a classic image interpolation algorithm; HAN is an SR method based on hierarchical attention; MHAN is an SR method that introduces multi-head attention; CTNet is an SR method combining convolution and Transformer; HSENet is an SR method based on efficient feature enhancement; MSGAN is an SR method based on generative adversarial networks; LKFN is an SR method based on large kernel convolution; TTST is a Transformer SR method based on Top-K selection; CATANet is an SR method that integrates channel attention and Transformer; and FMSR is an SR method based on multi-scale feature fusion.
[0078] Table 1 presents the comparison results under the conditions of reconstruction magnification of 2 and 4 using the above three evaluation indicators, where HR represents the original high-resolution remote sensing image baseline.
[0079] Table 1. Comparison of the present invention with nine other remote sensing image SR methods (AID)
[0080] As shown in Table 1, compared with the other nine remote sensing image SR methods, the present invention and other lightweight methods, HPD (Hierarchical Progressive Distillation) have achieved significant improvements in the reconstruction quality of high-resolution remote sensing images.
[0081] like Figure 7 As shown, a specific embodiment is used in this embodiment, which uses the RSCN7 dataset, which contains 2,800 images divided into 7 categories, each containing 400 images with a resolution of 400×400.
[0082] During the data processing phase, all training, validation, and test images were first downsampled to 256×256 resolution using bicubic interpolation as high-resolution (HR) reference images. These were then further downsampled at 2x and 4x scales to generate corresponding low-resolution (LR) images. The last 40 images for each category in the dataset were used as the test set, with the remainder used for training and validation. This data setup covers different terrain structures, comprehensively considers training requirements under various scenarios, and simulates the model's performance under diverse conditions.
[0083] The SR reconstruction results were evaluated using three metrics: peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and feature similarity, in order to verify the performance of SR reconstruction under the luminance channel.
[0084] During model training, the batch size was set to 32, the input image was randomly cropped into 32×32 patches, and common data augmentation strategies were employed, including random rotation and horizontal flipping. The optimizer used was the Adam optimizer, with parameters set to... , , .
[0085] In the ×2 super-resolution task, the model was trained from scratch for 500k iterations with an initial learning rate of 3×10⁻⁶. -3 The learning rate scheduling employs a cosine annealing restart strategy, with a period set to [200000, 200000, 100000], a restart weight of [1, 0.5, 0.3], and a minimum learning rate set to... .
[0086] In the ×4 super-resolution task, fine-tuning is performed based on the pre-trained ×2 model, which is also trained for 500k iterations with the same initial learning rate. At the same time, a restart strategy with a period of [200000, 200000, 100000] and a weight of [1, 0.5, 0.3] is used.
[0087] All experiments were implemented using PyTorch and were trained and tested on two NVIDIA GeForce RTX 4090 GPUs.
[0088] The remote sensing image SR methods selected for comparison include: Bicubic, HAN, MHAN, CTNet, HSENet, MSGAN, LKFN, TTST, CATANet, and FMSR. Bicubic is a classic image interpolation algorithm; HAN is an SR method based on hierarchical attention; MHAN is an SR method that introduces multi-head attention; CTNet is an SR method combining convolution and Transformer; HSENet is an SR method based on efficient feature enhancement; MSGAN is an SR method based on generative adversarial networks; LKFN is an SR method based on large kernel convolution; TTST is a Transformer SR method based on Top-K selection; CATANet is an SR method that integrates channel attention and Transformer; and FMSR is an SR method based on multi-scale feature fusion.
[0089] Table 2 presents the comparison results under the conditions of reconstruction magnification of 2 and 4 using the above three evaluation indicators, where HR represents the original high-resolution remote sensing image baseline.
[0090] Table 2 Comparison results of this invention with nine other remote sensing image SR methods (RSSCN7)
[0091] As shown in Table 2, compared with the other nine remote sensing image SR methods, the present invention and other lightweight methods, HPD (Hierarchical Progressive Distillation) have achieved significant improvements in the reconstruction quality of high-resolution remote sensing images.
[0092] According to the second aspect of this application, such as Figure 7 As shown, this embodiment provides a remote sensing image super-resolution system based on hierarchical progressive distillation, comprising: The first acquisition module is used to acquire raw remote sensing images and process the raw remote sensing images to obtain pseudo-high resolution images. The model building module is used to input the pseudo-high-resolution image into a pre-built hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high-resolution image reconstruction unit. The second acquisition module is used to encode the pseudo-high resolution image based on the shallow feature extraction unit to obtain shallow features; The third acquisition module is used to perform multi-stage feature processing on the shallow features in each of the hierarchical progressive distillation units, and to perform feature screening and information compression in the step-by-step processing to obtain initial deep features, and to perform spatial modeling and frequency domain enhancement processing on the initial deep features to obtain enhanced features. The fourth acquisition module is used to perform splicing and fusion processing on the enhanced features output by each of the hierarchical progressive distillation units based on the global feature fusion unit, and then fuse them with the shallow features through residual connection to obtain deep fusion features; The fifth acquisition module is used to perform convolution and pixel rearrangement processing on the deep fusion features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
[0093] Specifically, this embodiment corresponds one-to-one with the above method embodiments. The functions of each module have been described in detail in the corresponding method embodiments, so they will not be repeated here.
[0094] According to a third aspect of this application, this embodiment provides a computer-readable storage medium having a computer program stored thereon, the computer program including executable instructions that, when executed by a processor, implement the method described above.
[0095] The present invention can implement all or part of the processes in the above methods, or it can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or system capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0096] According to a fourth aspect of this application, a schematic diagram of the physical structure of an electronic device is provided, as shown below. Figure 8 As shown, the electronic device may include a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus. The processor can call logical instructions in the memory to execute a remote sensing image super-resolution method. This method includes: acquiring an original remote sensing image and processing the original remote sensing image to obtain a pseudo-high-resolution image; inputting the pseudo-high-resolution image into a pre-constructed hierarchical progressive distillation network, the hierarchical progressive distillation network including a shallow feature extraction unit, a hierarchical progressive distillation unit, a global feature fusion unit, and a high-resolution image reconstruction unit; encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features; performing multi-stage feature processing on the shallow features based on the hierarchical progressive distillation unit, performing feature filtering and information compression during the step-by-step processing to obtain initial deep features; performing spatial modeling and frequency domain enhancement processing on the initial deep features based on the global feature fusion unit to obtain enhanced features; and performing convolution and pixel rearrangement processing on the enhanced features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
[0097] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the remote sensing image super-resolution method provided by the above methods. The method includes: acquiring an original remote sensing image and processing the original remote sensing image to obtain a pseudo-high-resolution image; inputting the pseudo-high-resolution image into a pre-constructed hierarchical progressive distillation network, the hierarchical progressive distillation network including a shallow feature extraction unit, a hierarchical progressive distillation unit, a global feature fusion unit, and a high-resolution image reconstruction unit; encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features; performing multi-stage feature processing on the shallow features based on the hierarchical progressive distillation unit, performing feature filtering and information compression during the step-by-step processing to obtain initial deep features; performing spatial modeling and frequency domain enhancement processing on the initial deep features based on the global feature fusion unit to obtain enhanced features; and performing convolution and pixel rearrangement processing on the enhanced features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the remote sensing image super-resolution method provided by the above methods. The method includes: acquiring an original remote sensing image and processing the original remote sensing image to obtain a pseudo-high-resolution image; inputting the pseudo-high-resolution image into a pre-constructed hierarchical progressive distillation network, the hierarchical progressive distillation network including a shallow feature extraction unit, a hierarchical progressive distillation unit, a global feature fusion unit, and a high-resolution image reconstruction unit; encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features; performing multi-stage feature processing on the shallow features based on the hierarchical progressive distillation unit, performing feature filtering and information compression during the step-by-step processing to obtain initial deep features; performing spatial modeling and frequency domain enhancement processing on the initial deep features based on the global feature fusion unit to obtain enhanced features; and performing convolution and pixel rearrangement processing on the enhanced features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A remote sensing image super-resolution method based on hierarchical progressive distillation, characterized in that, include: Acquire raw remote sensing images and process them to obtain pseudo-high resolution images; The pseudo-high-resolution image is input into a pre-constructed hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high-resolution image reconstruction unit. Based on the shallow feature extraction unit, the pseudo-high resolution image is encoded to obtain shallow features; In each of the hierarchical progressive distillation units, the shallow features are subjected to multi-stage feature processing. During the step-by-step processing, feature selection and information compression are performed to obtain initial deep features. Furthermore, the initial deep features are subjected to spatial modeling and frequency domain enhancement processing to obtain enhanced features. Based on the global feature fusion unit, the enhanced features output by each of the hierarchical progressive distillation units are spliced and fused, and then fused with the shallow features through residual connections to obtain deep fused features; Based on the high-resolution image reconstruction unit, the deep fusion features are subjected to convolution and pixel rearrangement processing to obtain a high-resolution remote sensing image.
2. The method according to claim 1, characterized in that, The process of encoding the pseudo-high-resolution image based on the shallow feature extraction unit to obtain shallow features includes: A lightweight feature extraction method is used to encode the input image to obtain shallow features.
3. The method according to claim 1, characterized in that, The hierarchical progressive distillation unit includes a first hierarchical progressive distillation block, a second hierarchical progressive distillation block, and a third hierarchical progressive distillation block. It performs multi-stage feature processing on the shallow features, including feature filtering and information compression during the progressive processing to obtain initial deep features, including: The shallow features are input into the first hierarchical progressive distillation block, and the shallow features are processed based on the first channel attention to obtain the first distillation features. The shallow features are then processed based on the first partial large kernel convolution to obtain the first refined features. The first distillation feature and the first refining feature are input into the second hierarchical progressive distillation block. The first refining feature is processed based on the second channel attention to obtain the second distillation feature. The first refining feature is processed based on the second partial large kernel convolution to obtain the second refining feature. The second distillation feature and the second refining feature are input into the third hierarchical progressive distillation block. The second refining feature is processed based on the third channel attention to obtain the third distillation feature. The second refining feature is processed based on the third partial large kernel convolution to obtain the third refining feature. The fourth refining feature is obtained by processing the third refining feature through a single BSConv operation. The initial deep characteristics are obtained by fusing the first distillation characteristics, the second distillation characteristics, the third distillation characteristics, and the fourth refining characteristics.
4. The method according to claim 1, characterized in that, The hierarchical progressive distillation unit further includes a spatial-frequency domain hybrid attention enhancement block, which comprises a dynamic window attention subunit and a multi-scale frequency domain subunit. The spatial modeling and frequency domain enhancement processing of the initial deep features to obtain enhanced features includes: The initial deep features are input into the dynamic window attention subunit to obtain spatially enhanced features; The initial deep features are input into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features; The initial deep features, the spatial enhancement features, and the frequency domain enhancement features are self-learned and fused to obtain the initial enhancement features; The initial enhanced features are sequentially processed by convolutional projection pixel normalization, and then added to the shallow features through residual connections to obtain the final enhanced features.
5. The method according to claim 4, characterized in that, The step of inputting the initial deep features into a dynamic window attention subunit to obtain spatially enhanced features includes: The optimal window size for the shallow features is obtained based on a lightweight prediction network. Based on the optimal window size, the initial deep features are divided into several windows. Multi-head self-attention is calculated in each window to obtain spatially enhanced features.
6. The method according to claim 4, characterized in that, The step of inputting the initial deep features into a multi-scale frequency domain sub-unit to obtain frequency domain enhanced features includes: The first frequency domain feature corresponding to the initial deep feature is obtained based on the Fast Fourier Transform; The first frequency domain feature is subjected to downsampling, feature enhancement and upsampling operations at a predefined scale to obtain the second frequency domain feature corresponding to the predefined scale. After fusing the second frequency domain features corresponding to each scale, the first frequency domain features are added together, and then mapped back to the spatial domain through inverse fast Fourier transform to generate a spatial attention map. The initial deep features are processed based on the spatial attention map to obtain frequency domain enhanced features.
7. The method according to claim 1, characterized in that, The process, based on the global feature fusion unit, involves splicing and fusing the enhanced features output by each hierarchical progressive distillation unit, and then fusing them with the shallow features via residual connections to obtain deep fused features, including: The enhancement features are gradually refined through several hierarchical progressive distillation blocks to obtain several refined features; Several refined features are spliced and fused together to obtain deep fused features; The shallow features and the deep fusion features are aggregated through residual connections to obtain deep fusion features, which are then reconstructed into a high-resolution remote sensing image through pixel rearrangement operations.
8. A remote sensing image super-resolution system based on hierarchical progressive distillation, characterized in that, include: The first acquisition module is used to acquire raw remote sensing images and process the raw remote sensing images to obtain pseudo-high resolution images. The model building module is used to input the pseudo-high-resolution image into a pre-built hierarchical progressive distillation network, which includes a shallow feature extraction unit, several hierarchical progressive distillation units, a global feature fusion unit, and a high-resolution image reconstruction unit. The second acquisition module is used to encode the pseudo-high resolution image based on the shallow feature extraction unit to obtain shallow features; The third acquisition module is used to perform multi-stage feature processing on the shallow features in each of the hierarchical progressive distillation units, and to perform feature screening and information compression in the step-by-step processing to obtain initial deep features, and to perform spatial modeling and frequency domain enhancement processing on the initial deep features to obtain enhanced features. The fourth acquisition module is used to perform splicing and fusion processing on the enhanced features output by each of the hierarchical progressive distillation units based on the global feature fusion unit, and then fuse them with the shallow features through residual connection to obtain deep fusion features; The fifth acquisition module is used to perform convolution and pixel rearrangement processing on the deep fusion features based on the high-resolution image reconstruction unit to obtain a high-resolution remote sensing image.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program includes executable instructions that, when executed by a processor, implement the method of any one of claims 1-7.
10. An electronic device, characterized in that, include: One or more processors; A memory for storing executable instructions of the processor, which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1-7.